Papers with multimodal Large Language Models

5 papers
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness (2025.acl-long)

Copied to clipboard

Challenge: Existing models that assess mLLMs on harmful meme understanding are inaccurate and lack accuracy.
Approach: They propose a framework that adaptively probes the reasoning capabilities of mLLMs . their framework systematically reveals the varying performance of different target mllms a .
Outcome: The proposed framework systematically reveals the performance of different target mLLMs.
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for image captioning are primarily designed for short captions and are not suitable for long captions.
Approach: They propose an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework.
Outcome: The proposed metric outperforms existing metrics and achieves superhuman performance on LongCap-Arena.
Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data (2024.findings-acl)

Copied to clipboard

Challenge: OpenAI's GPT-4 has demonstrated remarkable multimodal capabilities, but specific mechanics of GPT4 remain unknown.
Approach: They propose a data collection methodology that synchronously synthesizes images and dialogues for visual instruction tuning.
Outcome: The proposed method improves on ten commonly assessed models and provides greater flexibility compared to existing methods.
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation approaches focus on mLLMs’ detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts.
Approach: They propose an agent-based arena-style evaluation framework that provides context-aware and unbiased assessment for mLLMs’ understanding of multimodal harmfulness.
Outcome: The proposed framework reduces evaluation biases of judge agents and provides unbiased comparisons of mLLMs’ abilities to interpret multimodal harmfulness.
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in multimodal Large Language Models (LLMs) have significantly enhanced the automation of medical image analysis, but still suffer from hallucinations and clinically significant errors.
Approach: They propose a grounding fixation strategy that integrates radiologist eye fixations and bounding box annotations into the LLM prompting framework.
Outcome: The proposed model improves performance without retraining across domain-specific and general-purpose models and achieves an 87.3% clinical average performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations